Tencent Unveils WorkBuddy Bench: A Coding Intelligent Agent Testing Ground Integrated with Code, Web, Office, and Security
Tencent has released the WorkBuddy Bench multi-domain evaluation suite, with a paper published on arXiv. It breaks the fragmented approach to evaluating coding intelligent agents and the lack of transparency in production benchmarks, integrating four types of work scenarios — repository-level code engineering, front-end artifacts, office automation — into one platform. The biggest highlight is not the volume of questions, but rather the design of the questions themselves, which prevents memorization of answers, ensuring that the evaluation truly reflects the generalization and transfer abilities of intelligent agents across different domains.